Metabarcoding and Metagenomics
● Pensoft Publishers
Preprints posted in the last 30 days, ranked by how well they match Metabarcoding and Metagenomics's content profile, based on 14 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Yepes Narvaez, V.; Rodriguez-Sanchez, A.; Atencia-Galindo, M. A.
Show abstract
The marine biodiversity inhabiting rocky shores in the Colombian Pacific remains largely undocumented, primarily due to geographic isolation, logistical challenges, and socio-political constraints. To address the existing knowledge gap, we conducted an expedition to enhance baseline biodiversity knowledge in rocky shores by integrating multiple complementary approaches, including visual censuses, specimen collection with morphological identification, environmental DNA (eDNA) metabarcoding and DNA barcodes. eDNA samples were collected at four coastal sites adjacent to rocky substrates, along with biological specimens obtained from fourteen locations through SCUBA diving at depths ranging from 1 to 25 meters. Tissue samples were subjected to genomic DNA isolation, followed by the generation and validation of cytochrome c oxidase subunit I (COI) barcode sequences, which were subsequently corroborated through taxonomic assessment to ensure accurate species identification. eDNA metabarcoding analyses yielded over 7 million high-quality sequence reads. Although taxonomic resolution at the species level was constrained by the limited completeness of reference sequence databases, a total of 106 species and 83 families were successfully identified, predominantly within the classes Actinopteri, Chondrichthyes, and marine mammals. From the 769 specimens obtained we generated 871 sequences, including 414 validated COI barcodes representing 76 species across 64 families. The integration of DNA barcoding and eDNA approaches resulted in over 1,400 taxonomic detections spanning five phyla, with only six species shared between methodologies. Richness and diversity varied among sites, and revealed significant differences along the coastline between Jurado and Cupica Gulf. All sequences were deposited in BOLDsystems database under the CCBIO project and were visualized through OBIS and GBIF databases. These findings provide the first molecular-based baseline for rocky shore biodiversity in the Colombian Pacific, highlighting the value of integrative approaches for monitoring and conservation.
Kirtane, A. A.; Weber, A. A.-T.
Show abstract
Passive sampling is the deployment of a collection material in the environment to continuously capture environmental DNA (eDNA) over time, offering the potential to integrate biodiversity signals while reducing the need for repeated active water collection. However, the mechanisms governing eDNA capture and retention on passive samplers remain poorly understood, limiting the interpretation of passive eDNA signals and their broader application. Here, we investigated the mechanistic performance of glass fibre passive samplers using controlled mesocosm experiments with three invasive freshwater bivalves: zebra mussels (Dreissena polymorpha), quagga mussels (Dreissena bugensis), and Asian clams (Corbicula fluminea). Specifically, we quantified eDNA accumulation dynamics, evaluated the contribution of different eDNA states, tested the persistence of captured eDNA, and compared passive sampler signals with conventional active sampling. Passive samplers rapidly accumulated target eDNA within hours of deployment, after which concentrations either plateaued or continued to increase depending on species. Sequential transfer of passive samplers between mesocosms containing different species showed that previously captured eDNA declined while new target eDNA accumulated to concentrations comparable to freshly deployed samplers, demonstrating continual turnover rather than permanent retention. Dissolved eDNA showed little evidence of accumulation beyond the concentration retained in the pore water within the membrane, suggesting that it is unlikely to be the dominant contributor to long-term passive sampler signals. Instead, the observed variability among replicate samplers, together with the physical properties of glass fibre membranes, suggests that membrane-bound and particulate eDNA are the primary contributors to passive eDNA capture. Collectively, these findings support a model in which glass fibre passive sampler signals reflect a dynamic equilibrium between ongoing eDNA capture and concurrent loss processes rather than cumulative accumulation over time. This mechanistic framework provides a foundation for interpreting passive eDNA data and informs the future development of passive sampling materials, deployment strategies, and biodiversity monitoring applications.
Feng, V.; Lin, H.-M.; Srivathsan, A.; Wang, H.; Lee, L.; Pedales, R.; Oberschmidt, D.; Meier, R.
Show abstract
1. Most species are neither discovered nor named, let alone included in analyses that require biological information such as trait measurements, images, ecological information and genome-scale data. Specimen-level DNA barcoding can help discover many of these species rapidly, but everything beyond discovery requires vouchers organized into putative species. Yet, existing barcoding workflows lack efficient techniques for voucher recovery, creating a post-barcoding bottleneck that limits the ability of converting barcoded specimens into biological knowledge. 2. Here we present a low-cost, open-source workflow consisting of two stages. The first safeguards barcoded specimens by separating them from DNA extracts and transferring them from microplates into ethanol-filled glass vials. The second converts the resulting voucher collection into a searchable physical resource by linking barcode-derived molecular Operational Taxonomic Unit (mOTU) assignments to vial positions and enabling specimens to be sorted into putative species either manually or automatically using a newly developed open-access robot (SORTER). 3. We evaluated the workflow using 2,024 insect specimens distributed across 21 96-well plates. For the first stage, DNA separation and specimen transfer required approximately 15 minutes per plate. For the second stage, MOTUmapper generated retrieval coordinates in a few seconds, after which the 2,024 vouchers belonging to the 452 putative species could be recovered manually in 5 days or with SORTER in 5 hours. Throughout both stages, specimen identities remained linked to barcode sequences, metadata and storage positions. 4. Vouchers are the Rosetta stones of biology because they connect different kinds of data to the same specimens. By safeguarding these vouchers and making them searchable, the workflow converts barcode projects from one-time molecular surveys into reusable resources for ecological and evolutionary research.
Muffett, K. M.; Sporre, M.; Miglietta, M. P.; Eytan, R.
Show abstract
Ranges of small benthic fauna are notoriously difficult to assess. In some of these cases, modern eDNA methods can shed light on species occurrence. Here we conduct an exploratory study on the fish eDNA recoverable from the gastrovascular cavities of the easy-to-sample pore water siphoning benthic invertebrate, Cassiopea, across six sites within the Florida Keys. Twenty-seven fish 12S identities were recovered from water samples, two from sediment samples, and seventeen from Cassiopea gut swabs. In total, thirty-two different species were identified from nineteen families, including one shark species (Ginglymostoma cirratum), and five species of cryptobenthic reef fishes (f: Gobiidae, Labrisomidae). Additionally, five species were identified from medusae samples that were not recovered in water or sediment samples. The species identities recovered may provide insight into the fish in direct proximity to Cassiopea assemblages, as well as indicate that Cassiopea may accrue disproportionate eDNA from cryptobenthic reef fish compared to surrounding environmental samples. The unorthodox sampling technique of using eDNA recovered from jellyfish stomachs yields another avenue for epibenthic community data acquisition.
Shedleur-Bourguignon, F.; Theriault, W. P.; Thibodeau, A.
Show abstract
Full-length 16S rRNA gene sequencing using Oxford Nanopore Technologies has emerged as a promising approach to improve species-level resolution in microbiota studies. However, the accuracy of taxonomic assignment remains highly dependent on the bioinformatics s used to process Nanopore long-read data. Therefore, the only way to ensure a good level of certainty in obtained results is to use positive controls in the form of mock communities in the experimental designs. In this study, we compared the performance of Epi2Me 16S (using Minimap2 or Kraken2) workflows provided by Oxford Nanopore Technologies and an EMU workflow for full-length 16S rRNA gene analysis. Using a commercial mock community sequenced across multiple Nanopore runs, taxonomic assignment accuracy and reproducibility was evaluated. Epi2Me-Kraken2 exhibited 18 % of incorrect genus-level assignments and failed to identify 3 species present in the mock community. While Epi2Me-Minimap2 achieved an excellent genus-level classification, reporting 9 % of sequences assigned to a genus not in the mock community, species-level assignments were inconsistent for several community members such as Listeria. In contrast, EMU provided accurate and consistent species-level taxonomic profiles, with all species correctly identified while keeping the number of genus absent from the mock community at 1.2%. ImportanceThese results highlight that Epi2Me integrated workflows are not the best option for specie-level taxonomic assignation. More importantly, this paper underscores the importance of routine inclusion of positive controls for microbiota studies, in the form of mock communities, as a critical safeguard for accurate data interpretation. Without the use of a mock community, a paper published would be at risk of reporting wrong observations and inaccurate conclusions.
Gueguen, L.-M.; Mathieu, A.; Perin, O.; Droit, A.
Show abstract
Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.
Craine, J. M.; Darcy, J. L.; Devitt, J.; Leopold, D.; Miller, G. W.; Ralson, M.; Schulte, N.; Fierer, N.
Show abstract
Freshwater bioassessment relies on assessing aquatic assemblages to infer ecological conditions, yet conventional surveys require extensive field sampling, specimen processing, and specialized taxonomic expertise. Existing environmental DNA (eDNA) methods have not yet provided a practical alternative to conventional macroinvertebrate assays in part because current approaches cannot feasibly recover broad taxonomic diversity at sufficient taxonomic resolution. Here, we evaluated targeted hybridization capture of mitochondrial cytochrome oxidase I (COI) target sequences as a unified molecular approach for cross-phylum freshwater bioassessment. Environmental DNA was collected at 18 sites along 63 km of Boulder Creek spanning nearly 1,500 m of elevation from forested headwaters to agricultural plains. COI targets were enriched using custom RNA bait panels designed to target regional freshwater arthropods, annelids, and molluscs. Hybridization capture increased recovery of COI sequences [~]1,760-fold relative to unenriched shotgun libraries, generating Folmer-region COI contigs that averaged [~]400 bp. Across the watershed, we recovered sequences for approximately 450 macroinvertebrate genera across 8 phyla. Detected macroinvertebrate richness averaged 56 genera per site and increased down Boulder Canyon before declining downstream of the city. Macroinvertebrate assemblage composition from hybridization capture paralleled patterns observed with past conventional bioassessment. These results demonstrate that targeted hybridization capture enables robust, cross-phylum detection of species used for freshwater bioassessment from environmental DNA.
Alarcon-Cruz, G.; Jacobs, S.; Baldwin, B. G.; Seltmann, K.; LeBuhn, G.
Show abstract
Understanding the spatial distribution of species and their patterns of endemism is necessary for establishing effective conservation priorities. Despite the vital pollination services bees provide, Californias native bee distribution patterns remain largely unexplored relative to plants and butterflies. We analyzed bee species richness and endemism across California and their concordance with plant distributions. Richness was high across areas of the California Floristic Province, including the Sierra Nevada, San Francisco Bay Area and Central Coast, South Coast Ranges, and the Transverse and Peninsular ranges. Bee endemism was more localized, concentrated in the San Joaquin Valley, eastern Sierra Nevada and adjacent Great Basin, Sierra Nevada foothills, and California deserts. Because richness and endemism appear to operate at different spatial scales and likely respond to different environmental drivers, effective conservation strategies must address both. Additionally, conservation plans that incorporate both plant and bee diversity are needed to achieve more comprehensive biodiversity protection.
Hellerich, C.; Klein, A.-M.; Garratt, M.; Mupepele, A.-C.; Fornoff, F.
Show abstract
Wildflower plantings are an important conservation measure for supporting wild bee diversity. They aim to enhance floral resources for nutrition, but do not explicitly consider that bees also require suitable nesting and overwintering resources. Wildflower plantings may provide nesting habitat for ground-nesting bees, but the effects of soil management, e.g. ploughing, on ground-nesting bees are hardly known. To study how ploughing and age of wildflower plantings affect ground-nesting bees, we sampled bees on ploughed and unploughed wildflower plantings aged 0-4 years. We used emergence traps to sample bees directly after emergence from the ground, allowing inference on nest numbers. We found that ploughing and wildflower planting age negatively affected overwintering bee nest numbers. Nesting peaked in the year of wildflower planting establishment and declined thereafter, indicating lower habitat suitability at later successional stages. Annual ploughing caused the greatest reduction in nest numbers (-72 %), providing evidence for an ecological trap.
Labbancz, J.; Dhingra, A.
Show abstract
Developments in Nanopore sequencing have enabled telomere to telomere genomic assembly as a routine technique in genomic research. Nanopore DNA sequencing for genomic assembly is typically performed on native DNA molecules, making it particularly sensitive to the quality of input DNA, with contaminating molecules limiting data yields and reducing read quality. As pangenome analysis gains interest, particularly in non-model plant species which are often rich in inhibitory secondary metabolites, the development of methods which can improve the quality and throughput of nanopore sequencing is essential. Here we describe a method for isolation of total DNA from the leaf tissues of diverse Viridiplantae species. The initial lysis buffer consists of a modified CTAB buffer, incorporating dimethyl sulfoxide for the reduction of viscosity, which can be problematic in many plant DNA preparations. An organic extraction with 2-butoxyethanol is utilized to further extract phenolic compounds which may be sufficiently hydrophilic to evade chloroform extraction, while reducing aqueous phase volume. Further cleanup via cesium chloride (CsCl) ultracentrifugation is performed to minimize the carryover of residual contaminating macromolecules. Samples prepared using this method are of consistent high quality, even when extracted from challenging late season leaf tissue or secondary metabolite rich species. Sequencing results from samples prepared by this method outperform those obtained from typical modified CTAB DNA isolation techniques in both quantity and quality. We tested sequencing performance from Vitis DNA isolated using a modified CTAB method and Vitis DNA isolated using the CsCl ultracentrifugation-based method described here. DNA isolated via the method described here produced 83% more >Q10 sequence data (52.61 Gb vs. 28.8 Gb), resulted in a 60% greater read N50 despite more handling steps (32.78kb vs. 20.45kb), and resulted in a higher modal read quality (Q27 vs. Q24). The consistency of this method across diverse plant taxa suggests its use as a general method for DNA isolation prior to Nanopore sequencing and genomic assembly for diverse plant taxa.
Xu, X.; Yang, X.
Show abstract
Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.
van Ooijen, R.; Buring, R.; Cornelius, A.; He, H.; van Oevelen, D.; Thieltges, D. W.; Hammoud, C.
Show abstract
The impact of invasive species on marine ecosystems is rapidly increasing, where they often outcompete native species in the absence of natural enemies. The parasite release hypothesis states that the success of invasive species relates partly to the loss of natural parasites during introduction and lower susceptibility to native parasites. Barnacles are highly successful invaders due to broad environmental tolerance and dispersal via shipping, but whether parasite release also participates in this success remains unknown. In this study, we analyse parasite infection patterns in native and invasive barnacles in the Wadden Sea by surveying communities across tidal zones. Additionally, year-round molecular monitoring of larval stages and a literature review were used to track the distribution of the invasive Pacific barnacle Balanus glandula in Europe and document its appearance in the Wadden Sea. The long-established invasive Austrominius modestus dominated the high and middle intertidal zone, whereas native species (Balanus crenatus and Amphibalanus improvisus) prevailed in lower zones. Native and invasive barnacles differed in parasite infection frequency (mostly cestodes and trematodes). The native Semibalanus balanoides had the highest prevalence (27%), followed by the invasive A. modestus (11%), and no infections were found in B. glandula. Lower parasite prevalence in invasive barnacles is consistent with the hypothesis that parasite release supports invasion success. In the absence of competent parasites, B. glandula could impact native barnacles through competition. Continued monitoring of B. glandula is recommended to track its distribution, interactions with native species, and parasite acquisition, providing further insight into the parasite release hypothesis.
Quijano, J. B.; Tayaban, K.; Baquiran, J. I. P.; Maala, G. J.; Requilme, J. N. C.; Sayco, S. L. G.; Dolorosa, R. G.; Cabaitan, P. C.; Conaco, C.
Show abstract
Giant clams are some of the largest bivalve molluscs. They form a vital partnership with Symbiodiniaceae dinoflagellates that supply most of their energetic requirements. However, the factors that shape giant clam-associated photosymbiont communities remain unknown. Here, we profiled Symbiodiniaceae communities using ITS2 metabarcoding in eight giant clam species (Hippopus hippopus, H. porcellanus, Tridacna crocea, T. derasa, T. gigas, T. maxima, T. noae and T. squamosa) from 11 sites across the Philippine archipelago. Symbiodiniaceae community structure was shaped by an interplay between giant clam host and environment. Most giant clams were dominated by members of a single symbiont genus, with Cladocopium as the most prevalent, followed by Durusdinium and Symbiodinium. However, giant clam hosts also exhibited flexibility in their symbiotic partners that was evident across sites. Differences in giant clam-associated symbiont communities may contribute to differences in holobiont function and adaptability to variable environments. These findings deepen our understanding of giant clam-Symbiodiniaceae associations, offering a framework for predicting how giant clams may be affected by increasingly stressful reef conditions and, more importantly, informing strategies to improve mariculture and conservation practices.
Monaghan, A. I. T.; Griffiths, N. P.; Sellers, G. S.; Lawson Handley, L.; Nunn, A. D.; Hänfling, B.; Macarthur, J. A.; Wright, R. M.; Cattaneo, M.; Bolland, J. D.
Show abstract
Context Pumping stations pose a threat to fish globally through land use change, habitat fragmentation and entrainment risk, with the catadromous and critically endangered European eel particularly impacted. Objectives/methods Establish, model, assess and understand the present-day distribution of European eel and resident fishes in 152 pumping station catchments in a once extensive wetland (The Fens) using eDNA metabarcoding (855 samples over two and half years), with specific focus on anthropogenic influences on hydrological connectivity and habitat quality. A removal survey design maximised confidence in negative results while minimising time and consumable costs. Results Eel occurrence upstream of pumping stations was low (occupancy = 28.3%) and positively associated with catchment area, fish species richness and natural hydrological connectivity (gravity drainage or flooding) and negatively associated with distance from the tidal limit. Fish species richness replaced catchment area and improved model performance, potentially acting as a biotic indicator of habitat quality and connectivity. Pumped catchments with manually operated upstream water transfers had reduced eel presence, potentially linked to the direction of water flow or the timing of operation. By contrast, fish species richness increased in these catchments during summer, suggesting displacement into unsuitable long-term habitats. Physical habitat maintenance had no detectable effect on eel occurrence or fish species richness. Conclusions This study provides the first landscape-scale assessment of European eel distribution and drivers of occurrence in pumped river catchments. The highly novel and comprehensive insights have implications for European eel conservation as well as infrastructure and catchment management, including compliance with legislation (EC Regulation No. 1100/2007).
Banos Lara, E.; Ras Segura, C.; de Boer, E. J.; Cundy, A. B.; Turon Barrera, X.; Nogue, S.; Holman, L. E.; Rius, M.
Show abstract
Replication is central to most experimental and sampling designs, increasing inferential power and capturing fine-scale data heterogeneity. However, its importance remains poorly evaluated in some ecological and evolutionary settings. This is the case of metabarcoding studies using DNA recovered from sedimentary archives, in which biological signals may integrate ecological information through depositional and burial processes, and are often inferred from a single sediment core per site. Here, we evaluated the effect of different types of replication using sedimentary DNA (sedaDNA) metabarcoding data from two genetic markers (mitochondrial COI and nuclear 18S), under a nested sampling design. The design included three intertidal sites, three spatially separated sediment cores per site (biological replicates), two sediment depth horizons per core, and eight PCR (technical) replicates per sediment sample. Variance partitioning showed that site identity and sediment age group together explained >70% of the variation in beta diversity, indicating that among-site spatial variation and stratigraphic variation were the dominant drivers of community composition. In contrast, variation among different cores within sites was small and non-significant (<5%). Among PCR replicates from the same sediment sample, richness varied substantially, whereas Shannon diversity was more consistent. Despite this variability, differences in community composition among technical replicates remained smaller than among biological replicates and site identity, indicating limited influence on broader ecological patterns. Community composition was highly similar among replicate cores within sites, consistent with stratigraphic coherence. These results indicate limited within-site heterogeneity and suggest that, under stratigraphically coherent conditions, increasing biological replication may yield limited additional information, whereas enhancing technical replication and stratigraphic resolution can improve ecological inference from sedaDNA metabarcoding datasets.
Tan, P.; Yadav, N.; Hauxwell, C.; Kerns, D. R.; Wilson, B.; Quinn, N.; Esquivel, I. L.; Rustgi, S.; Hernandez Europa, Y.; Patrick, D.; Ahmed, M. Z.
Show abstract
Heliococcus summervillei is an emerging invasive mealybug that causes severe dieback in grasses in pastures and turfgrass landscapes. It is widespread in Australia and has recently been detected across the Caribbean, Mexico, and the United States. Accurate identification of mealybugs is challenging due to cryptic morphology, overlapping diagnostic characters, and limited taxonomic expertise and literature, which makes molecular tools essential for regulatory diagnostics and management. We developed the first Cytochrome Oxidase I (COI) barcode for H. summervillei and used it to examine mitochondrial variation across available populations. COI sequences reveal approximately a 10.2% mitochondrial split between the Type A and Type B variants. Phylogenetic, haplotype network, and genetic distance analyses show that all invasive range populations share one haplotype associated with a recent invasion in the United States, Australia, Pakistan, and the Caribbean, whereas the Barbados lineage contains two closely related haplotypes that represent a historically stable mitochondrial variant. Together, these results establish the first COI reference library for H. summervillei, clarify mitochondrial lineage structure, and provide a practical barcode tool that enables rapid identification of invasive populations and supports timely regulatory and pest management responses. Recognizing mitochondrial variants also establishes a framework for resolving lineage-specific biological and management traits and strengthens reconstruction of introduction pathways central to regulatory decision-making and limiting further spread.
Xu, C.; Schalkwyk, H. V.; Powell, O.; Gustave, C.; Ball, L.; Ross, K.; Murray, E.; Aguirregoicoa, H.; Mackins, H.; Swinnerton, K.; Creedy, T. J.; Sivess, L.; Jones, J.; Castillo, K.; Bleet, R.; Salatino, S.; Mendis, Y.-T. C.; Lebre, P.; Mkrtchyan, H.; Cuber, P.
Show abstract
The reintroduction of extinct or endangered species to restore ecosystem function is an essential aspect of rewilding. The Wilder Blean Project at West Blean and Thornden Woods in Canterbury, UK, is committed to rewilding natural processes and enhancing biodiversity in one of England's oldest and largest areas of ancient woodland. The introduction of European bison (Bison bonasus) is an important part of the project. However, how the reintroduction of large herbivores influences local biodiversity and ecosystem functions during the early stages of rewilding remains poorly understood. Soil samples were collected from the same sampling sites before and two years after bison were reintroduced and profiled by metagenomic sequencing using Oxford Nanopore Technologies sequencing platforms. The results showed that the alpha diversity of soil organisms did not change significantly before and after the introduction of European bison, while beta diversity showed modest shifts in community composition. The relative abundance of some nitrogen-fixing and photosynthetic microbial genera showed declines in the 2024 Bison Area, while the mycorrhizal fungus genus Rhizophagus was significantly less abundant than in the 2024 Control Area. Despite relatively stable taxonomic diversity, functional composition differed significantly between the 2022 and 2024 Bison areas and among the 2024 rewilding treatments, revealing a decoupling between taxonomic diversity and functional composition. Amino acid synthesis pathways and carbon metabolism pathways were significantly enriched. These findings highlight the potential of long-read Oxford Nanopore metagenomics to reveal functional shifts that may not be apparent from taxonomic diversity alone. Although these early-stage responses cannot yet predict long-term rewilding trajectories, continued longitudinal monitoring integrating microbial, soil physicochemical, and ecosystem-level measurements will be essential to determine the persistence and ecological significance of these functional shifts.
Bibi, A.; Iqbal, T.; Ilyas, K.; Nosheen, A.
Show abstract
The Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) and associated nuclease gene (Cas), originating from the bacteria acquired immune system, have revolutionized gene editing technology. In this regard, type II (Cas9) been extensively studied and widely applied CRISPR system so far. The mechanism for precise manipulation of genomic sequences is guided by small RNA called CRISPR RNA (crRNA). In this study we devised and optimized CRISPR-Cas9 screening system based on Cas9 gene detection, targeting a conserved part of recognition domain (REC) consisting of arginine rich bridge helix (BH). We used hemi-nested PCR approach for screening sensitivity and reproducibility. The recombinant E. coli DH5 alpha containing the pRGEB32 vector (DH5 alpha/pRGEB32) with the Cas9 gene was used for system optimization. Subsequently, the screening system was applied and validated on different environmental bacterial strains including Alcaligenes faecalis and Pseudomonas stutzeri, isolated from sewerage samples. The optimized hemi-nested PCR resulted in amplification of targeted region in environmental bacterial strains and results were reproduced successfully. Furthermore, nucleotides and amino acid sequence, motif and domain analysis of PCR products, confirmed the targeted Cas9 REC-BH domain. Presently, no rapid and cost effective CRISPR-Cas screening system is available except expensive whole genome sequencing approach. Our investigation aimed to device rapid and cost effective screening system for identification of new variants of Cas9 proteins in environmental bacterial species. In this context, the developed Cas9 gene-based CRISPR-Cas screening system (C9CSS) may be a potential rapid screening tool to identify new Cas9 orthologs in different bacterial genomes with improved functions.
Pradhan, P.
Show abstract
Global Biodiversity Information Facility (GBIF) occurrence retrievals for an irregularly shaped region are limited by the API spatial query capabilities - rectangular envelopes or size/vertex-limited WKT polygons - neither of which conform to protected areas, sacred groves, wetlands, panchayat or municipal boundaries or any other arbitrary KML polygon of interest queried by users. This paper presents and validates an open, self-contained, adaptive spatial-tiling protocol that (i) ingests any KML polygon of any shape, size and location on earth, breaks it into a set of GBIF API-compatible rectangular tiles, (ii) queries, cleans and clips the individual records to the target polygon, and (iii) summarises the inventory with a generic diversity-completeness-rarefaction module, with minimal manual re-parameterisation between sites. The protocol implements an iterative quadtree refinement algorithm that adapts tile number, size and location to the target polygon geometry, is combined with a fault-tolerant pagination/retry query system, a boundary-exact two-step clipping procedure and a Chao1-based completeness assessment to ensure statistical comparability between sites of different spatial extent and sampling intensity. The algorithm is implemented in open R source (sf, terra, rgbif, tidyverse) with the tiling algorithm controlled by the four parameters only (initial cell size, area floor, tile overlap threshold, recursion limit), with default settings on a new site by simply changing the input file path. This paper describes in detail its five main components - (i) polygon input and validation, (ii) quadtree adaptive tiling, (iii) polygon coverage verification, (iv) tile-wise GBIF query with retry/shrink pagination and partial data retention, (v) boundary-exact deduplication, clipping and diversity estimation. A downstream generic module estimates diversity, Chao1 richness/completeness and Hurlbert rarefaction, for each taxonomic rank and generates rank-ordered diversity tables as output. The generalisability of algorithm to multiple sites has been demonstrated with second polygon (Sonamukhi Sal forest dominated stretch, Bankura district, West Bengal; approx. 610 sq km) that differs from the first (Bishnupur Sal forest dominated stretch; 938 sq km) in both size and complexity (10 vs 34 KML vertices) and report the tiling and diversity metrics comparable results across the two polygons. With no parameter changes, the algorithm generated 135 adaptive query tiles for Sal forest dominated stretch adjoining Bishnupur, and 86 tiles for Sal forest dominated stretch Sonamukhi SDFP, covering completely the area of both polygons. The number of tiles per 100 sq km is comparable between the two runs (14.4 vs 14.1 tiles) despite the 35% difference in polygon size and 3.4x vertex count. The tile-wise querying with retry/shrink pagination retrieved 6,169 GBIF records (excluding errors) with boundary-exact clipping across 404 species for Bishnupur and 1,222 GBIF records (excluding errors) across 271 species for Sonamukhi; the generic diversity module processed the records without further parameter changes and generated comparable metrics for each rank at both sites. The protocol addresses a general bioinformatic challenge in polygon-based GBIF queries, is provided as an open, reusable, documented method which has been validated on two sites. Because the protocol has so far been validated on only two polygons that differ markedly in size, shape and observer regime, it may be regarded as an initial cross-site validation rather than a comprehensive benchmark, and recommend testing on a broader, globally distributed set of polygons before the approach is treated as a general-purpose standard.
Lewis, R.; Dodd, K.; Macadam, C.; Matthews, I.; Pulver, S. R.
Show abstract
Freshwater ecosystems in Scotland are increasingly threatened by multiple interacting stressors, including pollution, hydrological alteration, land use change, and climate-driven pressures. Scotland's rivers are amongst the most physically diverse and dynamic in the UK, contributing significantly to the country-s economy and natural heritage. Monitoring 125,000 km of waterways presents an issue, with limited funding and capacity of environmental agencies creating an environmental monitoring deficit at a time when robust data are essential for delivering national and global biodiversity targets. Citizen science offers a scalable, community driven approach to strengthening environmental evidence. This study evaluates the development, implementation, and outcomes of Guardians of our Rivers (GooR), a nationwide citizen science programme using macroinvertebrate-based river health assessments based on the Anglers' Riverfly Monitoring Initiative (RMI) established by the Riverfly Partnership. Between 2022 and 2025, GooR trained more than 850 volunteers, established 123 monitoring sites across Scotland, and generated over 860 surveys-surpassing the previous 16 years of aquatic invertebrate monitoring efforts. Standardised field training, full equipment provision, and structured support enabled high-quality data collection across diverse catchments. The establishment of nationally applied trigger and target thresholds created a consistent mechanism for detecting ecological deterioration and ensured that verified trigger breaches (n = 39 in 2025) resulted in formal investigation or follow-up. Analysis of national abundance patterns revealed ecological and biogeographic trends across the eight RMI taxa, confirming data sensitivity to habitat differences, water quality, and regional pressures. A detailed case study of the Lothian Esk catchment demonstrated the programme's capacity to detect seasonal macroinvertebrate dynamics and site level ecological trajectories over time. Overall, GooR illustrates how well-designed and well-supported citizen science can deliver high resolution spatiotemporal datasets, enhance early warning systems, empower communities, and make a meaningful contribution to promoting river health. Importantly, GooR forged a working partnership with Scottish Environmental Protection Agency (SEPA) that allowed communities to not only gather data but also transformed the way citizen science is recognized as contributing to freshwater conservation. Continued investment and long-term support will be essential to sustain these gains and realise the full potential of citizen-driven freshwater stewardship.